01 The Big Picture
Every frontier model spec sheet now reads like a split personality: "671B parameters" in the headline, "37B active parameters" in the footnote. That footnote is the whole story of modern LLM engineering.
Dense models wire every parameter into every token — capacity and per-token cost are the same number, growing together until serving becomes impossible (doc 08's memory wall). Mixture-of-Experts (MoE) breaks that coupling: store a huge parameter pool, but route each token through only a small, learned subset. The result — publicly reported across the current frontier — is models with dense-class latency and an order of magnitude more knowledge capacity.
This doc is the survey behind the concepts: the gate math, the load-balancing tricks, the capacity factor, the memory budget — and a model-by-model read of what Mixtral, DeepSeek-V3, GLM-4.5, Qwen3, Llama 4 and GPT-OSS actually shipped, as publicly reported.
02 What MoE Is — Precisely
A MoE layer replaces one feed-forward network (the FFN, which is ~2/3 of a transformer's parameters — doc 02) with N expert FFNs plus a router. For each token, the router scores all experts, picks the top-k, and blends only those outputs. Attention layers stay dense and shared by everyone.
Coarse MoE — Mixtral 8×7B
The design that made open MoE credible: 8 large experts (each ~7B params), top-2 routing. Total ≈ 47B parameters, but each token touches only ~13B — two experts plus the shared dense parts. Fine-grained it is not: with 8 experts and top-2 there are only C(8,2) = 28 possible expert combinations per layer.
Fine-grained MoE — the DeepSeek lineage
Instead of few large experts, use many small ones: DeepSeek-V3 (as publicly reported) splits its FFN into 256 routed experts, picks top-8, plus 1 shared expert per token. Combinatorially huge routing space → far more precise specialization per token, at the same active-parameter budget.
Why did fine-grained beat the 8-expert design? Combinations. With 8 experts top-2, a token's "sentence" of expert choices is 2 letters from an 8-letter alphabet. With 256 experts top-8, it's 8 letters from a 256-letter alphabet — the router can compose fine-grained capabilities (syntax + domain + format) instead of choosing between a few monolithic personalities. Same bytes read per token, much richer conditioning.
03 Why Sparsity Wins — Decoupling Capacity from Compute
The entire economic argument of MoE is one decoupling:
04 How It Works — One Token Through the Router
Step through a single token hitting one MoE layer. Watch the gate score every expert, keep only the top-k, and blend.
Every MoE layer in the model runs this same ceremony, per token, per layer. The router is itself a tiny learned matrix (W_g) — its only job is to spend the compute budget wisely. Everything else in this doc is about what goes wrong when 30 billion tokens all want the same expert at once.
05 The Gate Math
Four formulas run the entire MoE economy. All symbols: T = tokens in batch, N = number of experts, k = top-k.
Why the auxiliary loss was painful: L_aux pushes routing toward uniformity, but uniform routing is not what the model wants to learn — experts should specialize unevenly. The gradient conflict between "balance for hardware" and "specialize for quality" directly cost model performance. DeepSeek's reported fix decouples the two: the bias term bᵢ steers selection to keep experts busy while the actual gate values — what gets blended into the output — stay pure softmax of the true affinities. Balance becomes a control loop, not a gradient penalty.
Why capacity α ≈ 1.25: perfectly balanced routing would send exactly T·k/N tokens to each expert; the 1.25 headroom absorbs imbalance so a modestly popular expert doesn't overflow. It's a bargaining point: higher α wastes compute and buffer memory on idle slots; lower α drops more tokens (each dropped token's knowledge contribution for that layer is lost). Modern deployments without hard capacity limits instead penalize overflow softly — but the C formula is the vocabulary in every MoE paper.
Shared-expert regularization: the shared expert (always fired, every token) is told: "absorb the common denominator." Routing noise — trivia every token needs, grammar, generic transformations — no longer needs to be competed for in the top-k. That keeps routed experts free to specialize, and it's why the pattern (1 shared + k routed) recurs across DeepSeek-V3, GLM-4.5 and Llama 4 (as publicly reported).
06 Model-by-Model — The Frontier Spec Sheet
All figures below are publicly reported values as of writing; treat them as design landmarks, not contracts.
| Model | Params total → active | Experts / top-k | Attention | Context | Engineering fit |
|---|---|---|---|---|---|
| Mixtral 8×7B | 47B → 13B | 8 large / top-2 | GQA + sliding window | 32K | The existence proof: open MoE matching a much larger dense model at 13B-class latency. Coarse experts — 28 possible combinations per layer. |
| DeepSeek-V3 | 671B → 37B | 256 fine-grained / top-8 + 1 shared | MLA (KV compression — doc 22) | 128K | Frontier quality at ~dense-37B serving cost. Aux-loss-free balancing, multi-token prediction (MTP) head for speculative-style decoding. |
| GLM-4.5 | 355B → 32B | 160 / top-8 + 1 shared (reported) | GQA | 128K | Tuned explicitly for agentic coding + tool use; a smaller "Air" sibling (106B/12B) for cheaper serving. |
| Qwen3-235B-A22B | 235B → 22B | 128 / top-8 (reported) | GQA + QK-Norm | 128K | Open-weight flagship with mature tooling; the "A22B" naming convention itself encodes total-vs-active. |
| Llama 4 (Scout / Maverick) | 109B / 400B → 17B | 16 / 128 routed + shared; alternating dense–MoE layers | iRoPE (interleaved) for very long context | up to 10M claimed (Scout) | MoE only on some layers — a hybrid budget; native multimodal, huge-context ambitions. |
| GPT-OSS (120B / 20B) | 117B / 21B → 5.1B / 3.6B | 128 / top-4 (reported) | attention sinks | 128K | Open-weights MoE small enough for a single consumer GPU; sinks stabilize attention over long generations. |
07 Engineering Takeaways — Serving the Beast
Size your serving fleet by total params (memory) and budget speed by active params (bandwidth). Prefer models with a shared expert + aux-loss-free balancing for stable EP utilization. Watch expert-balance telemetry like you watch GPU utilization.
Don't assume "671B model" means 671B-slow — or 13B cheap to host. Don't plan single-GPU serving for large MoE. Don't compare models on total params alone: a 235B-A22B and a 355B-32B are in the same latency class.
08 Mental Models
A hospital employs hundreds of specialists (total params) but a patient only consults two or three per visit (active params). The payroll — memory footprint — is driven by the full staff; visit cost by the consultants seen. Triage (the router) must also prevent everyone queueing for the same cardiologist — that's load balancing.
Total params = the size of the library's collection; active params = how many shelves the librarian physically walks to per question. You can grow the collection indefinitely while keeping the walk short — but the building (HBM) must still house every book, and the librarian's route (EP dispatch) must avoid crowds.
Think of each expert as a specialized kernel and the router as a JIT dispatcher: per input, a few kernels are selected and fused (weighted sum), while the rest stay cold in memory. Fine-grained MoE = many small kernels + rich dispatch; capacity factor = guard bands in your thread pool scheduling.
09 Common Misconceptions
"A 671B MoE is 671B worth of quality in every token." No token ever touches more than ~37B of weights. Capacity is stored knowledge available across the routing space; per-token depth is the active budget. Quality gains are real but bounded by routing precision.
"MoE is 8× cheaper to serve than a dense 671B." It's ~18× cheaper in bandwidth and FLOPs per token, but you still buy ~0.7 TB of HBM per replica and pay all-to-all communication plus balancing overhead. Cheaper per token, not cheap to host.
"More experts is automatically better." More, smaller experts only help if the router learns to compose them; they add dispatch overhead and make load balancing harder (256 experts × top-8 is a much harder scheduling problem than 8 × top-2). The win came from fine-grained structure plus shared-expert isolation plus better balancing — together.
"Dropped tokens (capacity overflow) are a rounding error." At small batch sizes overflow is rare; at large EP deployments a bad α or a hot expert can silently drop a meaningful token fraction — quality loss that never shows up as an error, only as slightly worse outputs.